Papers with code-mixing of Hindi-English
MUTANT: A Multi-sentential Code-mixed Hinglish Dataset (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods to identify code-mixed text are difficult to scale effectively and efficiently on multi-sentential data. |
| Approach: | They propose to identify multi-sentential code-mixed text (MCT) from multilingual articles using a token-level language-aware pipeline. |
| Outcome: | The proposed dataset includes 67k articles with 85k identified Hinglish MCTs. |